Papers with dataset construction

13 papers
Efficient Online Scalar Annotation with Bounded Support (P18-1)

Copied to clipboard

Challenge: Existing methods for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments are not shown.
Approach: They propose a method for efficiently eliciting scalar annotations by human judgments.
Outcome: The proposed method leads to increased correlation with ground truth, suggesting it is an improved mechanism for dataset creation and manual system evaluation.
RethinkingTMSC: An Empirical Study for Target-Oriented Multimodal Sentiment Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that current TMSC systems rely on textual information, and the progress in tackling this task has slowed down.
Approach: They propose to integrate both visual and textual information to improve the performance of TMSC by considering multimodal information.
Outcome: The proposed model integrates both visual and textual information to improve performance.
BlendX: Complex Multi-Intent Detection with Blended Patterns (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets such as MixATIS and MixSNIPS have limitations in their formulation.
Approach: They propose a set of multi-intent detection datasets that feature more diverse patterns than their predecessors.
Outcome: The proposed datasets feature more diverse patterns than their predecessors and are more complex and diverse than existing datasets.
GLoHBCD: A Naturalistic German Dataset for Language of Health Behaviour Change on Online Support Forums (2022.lrec-1)

Copied to clipboard

Challenge: Existing motivational interviewing methods lack the deep understanding of user utterances that is essential to the spirit of motivational interviews.
Approach: They propose to use a German dataset of naturalistic language around health behaviour change to examine the motivational state of the user.
Outcome: The proposed dataset of naturalistic language around health behaviour change is based on a weight loss forum in germany and is evaluated using theoretically grounded motivational interviewing categories.
Pula: Training Large Language Models for Setswana (2025.naacl-long)

Copied to clipboard

Challenge: Setswana is a Bantu language spoken by an estimated five to ten million people worldwide.
Approach: They propose to make setswana-based models available for the first time using data available from setswa and setswegian databases.
Outcome: The proposed models outperform GPT-4o and Gemini 1.5 Pro on English-Setswana translation tasks and achieve state-of-the-art performance on Setswanan reasoning tasks.
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: Almost all popular summarization datasets do not come with inherent quality assurance guarantees.
Approach: They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data.
Outcome: The proposed metrics can be inexpensive heuristics for detecting generically low quality examples.
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh (2025.acl-long)

Copied to clipboard

Challenge: Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains.
Approach: They propose to open-source a large-scale instruction-following dataset covering key institutional and cultural knowledge relevant to Kazakhstan.
Outcome: The proposed dataset improves LLMs’ understanding of procedural, legal, and structural governance topics.
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur’anic Tafsir (2025.emnlp-main)

Copied to clipboard

Challenge: An annotated Tafsir ontology and a collection of 15 structured Tafsian books are presented in this paper.
Approach: They propose a framework for retrieval and question-answering Tafsir data that spans the entire pipeline from dataset construction through evaluation and error analysis.
Outcome: The proposed framework achieves 69.52% accuracy and 74.36% correctness overall, though multi-hop and context-dependent questions remain challenging.
Personality Understanding of Fictional Characters during Book Reading (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to predict characters' personalities have not been studied in the NLP field due to the lack of appropriate datasets mimicking the process of book reading.
Approach: They propose a dataset to predict characters' personalities that uses an exhaustive vocabulary of personality traits as targets.
Outcome: The proposed dataset is efficient and accurate and relies on long-term context to achieve accurate predictions for both machines and humans.
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios (2025.emnlp-main)

Copied to clipboard

Challenge: a number of tools are used to perform complex tasks, but the tool utilization process can cause errors.
Approach: They propose a critique evaluation benchmark for tool learning that analyzes function-calling errors on tool evaluation benchmarks.
Outcome: The proposed critique evaluation benchmark holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios.
Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm.
Approach: They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
Outcome: The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
Batayan: A Filipino NLP benchmark for evaluating Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages.
Approach: They propose a benchmark that systematically evaluates LLMs across three key natural language processing competencies: understanding, reasoning, and generation.
Outcome: The proposed benchmark covers eight tasks covering Tagalog and code-switched Taglish utterances.
ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to adapt image-focused models for video understanding have not been successful in analyzing long video sequences.
Approach: They propose a video instruction dataset that outperforms existing video instruction data for fine-tuning MLLMs by incrementally increasing input context length.
Outcome: The proposed model outperforms existing models on video benchmarks and outperformed proprietary models on VideoMME even with a compact 7B model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations